Repository navigation
fix: prevent orphaned crawl record when lock check fails in StartCrawler - #47
Merged
StJudeWasHere merged 4 commits intoMay 23, 2026
Conversation
SaveCrawl was called before addCrawler checked the in-memory lock, so a concurrent or rapid trigger could create a DB crawl row with a NULL end timestamp and then return an error — leaving that row permanently stuck and blocking all future crawls with "project is already being crawled". Move addCrawler before SaveCrawl so no DB record is written when the lock is already held. If SaveCrawl then fails, removeCrawler cleans up the lock entry so the caller can retry cleanly. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Move GetLastCrawl to after addCrawler so a rejected duplicate trigger skips the DB round-trip entirely — the result is only needed inside the goroutine which never starts when the lock is already held. Add TestStartCrawlerNoDuplicateDBRecord to assert that a second StartCrawler call while a crawl is in progress returns an error and does not invoke SaveCrawl, preventing the orphaned NULL-end-timestamp that blocked all future crawls. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
The type=sha,prefix={{branch}}- rule produces a tag like '-d6926f4'
on PR events because {{branch}} evaluates to empty — Docker rejects
tags that start with a dash. Restrict the sha tag to the default
branch only, where {{branch}} is always 'main'.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Fork PRs cannot write to the upstream registry — the GITHUB_TOKEN is read-only for packages in that context. Build the image on PRs to keep Dockerfile validation, but only push on branch push events. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Owner
|
Good catch. Thank you! |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
When two crawl requests arrive in rapid succession (or a scheduler fires
slightly early),
StartCrawlercould leave a zombie crawl row in thedatabase with
end = NULL, permanently blocking all future crawls with"project is already being crawled" — even after the real crawl finished.
Root cause
SaveCrawlwrites the DB record beforeaddCrawlerchecks thein-memory
crawlersmap. WhenaddCrawlerreturns an error, the functionexits early and the goroutine that would call
UpdateCrawl(which setsend) is never started — leaving the row stuck forever.Fix
Acquire the in-memory lock first, then do all DB writes.
GetLastCrawlisalso moved after the lock check since its result is only used inside the
goroutine — no point querying the DB on a request that will be rejected.
Changes
internal/services/crawler.go— reorderaddCrawlerbeforeGetLastCrawland
SaveCrawl; release the lock ifSaveCrawlfails.internal/services/crawler_test.go— new testTestStartCrawlerNoDuplicateDBRecordasserts that a second concurrentStartCrawlercall returns an error and does not invokeSaveCrawl.Test plan
go test ./internal/services/...passes including the new test.should be rejected and no new row with
NULL endappears incrawls.endtimestamp set.SaveCrawlfails (e.g. DB down), no orphaned in-memory lock entryremains and the project can be crawled once the DB recovers.
🤖 Generated with Claude Code